PROBER: Ad-Hoc Debugging of Extraction and Integration Pipelines
نویسندگان
چکیده
Complex information extraction (IE) pipelines assembled by plumbing together off-the-shelf operators, specially customized operators, and operators re-used from other text processing pipelines are becoming an integral component of most text processing frameworks. A critical task faced by the IE pipeline user is to run a post-mortem analysis on the output. Due to the diverse nature of extraction operators (often implemented by independent groups), it is time consuming and error-prone to describe operator semantics formally or operationally to a provenance system. We introduce the first system that helps IE users analyze pipeline semantics and infer provenance interactively while debugging. This allows the effort to be proportional to the need, and to focus on the portions of the pipeline under the greatest suspicion. We present a generic debugger for running post-execution analysis of any IE pipeline consisting of arbitrary types of operators. We propose an effective provenance model for IE pipelines which captures a variety of operator types, ranging from those for which full or no specifications are available. We present a suite of algorithms to effectively build provenance and facilitate debugging. Finally, we present an extensive experimental study on large-scale real-world extractions from an index of ∼500 million Web documents.
منابع مشابه
Survey the Security Function of Integration of vehicular ad hoc Networks with Software-defiend Networks
In recent years, Vehicular Ad Hoc Networks (VANETs) have emerged as one of the most active areas in the field of technology to provide a wide range of services, including road safety, passenger's safety, amusement facilities for passengers and emergency facilities. Due to the lack of flexibility, complexity and high dynamic network topology, the development and management of current Vehicular A...
متن کاملPipelines for Ad-hoc Large-scale Text Mining
Pipelines for Ad-hoc Large-scale Text Mining Today’s web search and big data analytics applications aim to address information needs (typically given in the form of search queries) ad-hoc on large numbers of texts. In order to directly return relevant information instead of only returning potentially relevant texts, these applications have begun to employ text mining. The term text mining cover...
متن کاملPipelines für effiziente und robuste Ad-hoc-Textanalyse
Suchmaschinen und Big-Data-Analytics-Anwendungen zielen darauf ab, ad-hoc relevante Informationen zu Anfragen zu finden. Häufig müssen dafür große Mengen natürlichsprachiger Texte verarbeitet werden. Um nicht nur potentiell relevante Texte, sondern direkt relevante Informationen zu ermitteln, werden Texte zunehmend tiefer analysiert. Dafür können theoretisch komplexe Pipelines mit zahlreichen A...
متن کاملA Visual Interface for on-the-fly Biological Database Integration and Workflow Design Using VizBuilder
Data integration plays a major role in modern Life Sciences research primarily because required resources are geographically distributed across continents and experts depend upon leveraging these digitally archived resources. The ever changing and exponentially growing digital archives pose a significant challenge for traditional data integration efforts and efficient development of data proces...
متن کاملIFAS: Interactive flexible ad hoc simulator
Ad Hoc networks are characterized by fast dynamic changes in the topology of the network. Introduction of new routing algorithms requires a deep and trustworthy evaluation process. In this paper we describe an Interactive Flexible Ad Hoc Simulator (IFAS) that presents a modern and novel approach to the family of Ad Hoc simulators. This simulator supports unique viewing, debugging, tuning and in...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
- CoRR
دوره abs/1004.1614 شماره
صفحات -
تاریخ انتشار 2010